Repository navigation
feat(diagrams): the agent draws in the chat — render_diagram, phase 1 of the Rich Diagrams spec - #104
Merged
Merged
Conversation
…fore changing it
The chat's Markdown renderer lives inline in media/chat.html, and no suite ran
it. The next commit changes how it finds a code fence, so this one writes down
what it does today, against today's code:
an ordinary fenced block a code block; the language line is not in it
two blocks, prose between and a block still being typed is shown as code
inline code escaped, never formatted; a lone backtick is text
a path in inline code a file chip, when it names a real project file
a fence inside a list item still a code block, indented as it was
a fence stuck to a line "Run this: ```bash" opens a block, and a closing
fence stuck to the last line of code closes one
— sloppy, and already tolerated
The functions are sliced out of chat.html itself, the way shHighlight.test.js
does it, so the suite runs the shipped code and not a copy of it.
…aragraph came out a fragment per line
An answer that mentioned a code fence in passing — "each closed by a plain
```` ``` ````" — was rendered as two code blocks holding one backtick each, and
the rest of the paragraph arrived one streamed fragment per line: "gr",
"ounded ent", "irely in the actual code". Reported from a real session. The
model had quoted the fence correctly, in an inline span of four backticks.
Two things were wrong, and the second is what made it look so bad.
render() split the text on every run of three backticks, wherever it stood. So
```` ``` ```` was three fences: a block containing "` ", and then an unfinished
block containing the rest of the message.
lastStableIndex() decides what the streaming renderer may freeze — move out of
the live tail and into the page for good. It counted fences the same way, and
when the text after an even-numbered one had no newline yet, it answered "all
of it". From then on every delta was frozen as it arrived, each in its own
<div>.
A fence is a line. mdSegments() reads the text line by line: a block opens on a
line that is, after its indentation, three or more backticks and a language (or
anything else without a backtick in it), and closes on a line ending in a run
at least as long as the one that opened it. render() and lastStableIndex() both
use it, so they cannot disagree again. A block counts as finished only once its
closing line has its newline; text after it stays live until the next block
finishes.
Inline code takes the count seriously too: a run of N backticks opens a span
that the next run of exactly N closes, on the same line. That is how three
backticks are quoted (inside four), and one (inside two).
Kept, because models do it and the old code let them: a fence stuck to the end
of a line of prose ("Run this: ```bash") still opens a block — when nothing but
a one-word language follows it and it is not the closing half of an inline
span — and a closing fence stuck to the last line of code still closes one.
Not changed: tildes are not fences; and a line of prose that ENDS in three bare
backticks still opens a block, as it did. CommonMark would not; a model that
wants to say "```" has an inline span for it, and now it works.
chatMarkdownFences.test.js gains the report itself, whole and streamed — in
sixty different chunkings, and one character at a time; the property that
makes streaming safe (streamed and whole are the same page, and only finished
blocks are frozen) over six documents and 400 random mixes of backticks,
newlines and words; N-backtick spans; fences nested by length; language lines
that say more than a word; and that nothing in a message becomes markup. The
six pins from the last commit pass unchanged. The ten new tests fail on the old
code.
Replayed in Chromium against the shipped page with the session's exact text, in
184 deltas: the paragraph was 57 lines and two code blocks before, and is one
paragraph and none after. 21 single-edit mutations of the new code are each
caught by the suite.
…rror at once, and the repair ladder
First slice of rich diagrams; docs/RICH-DIAGRAMS.md arrives with the last one.
The agent draws flows and architectures out of box-drawing characters. The plan
is that it describes STRUCTURE instead — nodes, edges, groups, one accent,
never a coordinate or a colour — and the editor draws. This commit is the part
that decides whether a description may be drawn. Nothing calls it yet.
diagram/schema.js the versioned schema, and a small interpreter for the
subset of JSON Schema it uses. One object is both the
tool's input schema and what a spec is checked against.
Two tiers of limits: HOUSE is what a model is held to
(12 nodes; labels 28, second lines 32, edge labels 20,
title 80; groups two deep), HARD is what the renderer
will still draw (24 nodes, 48 edges).
diagram/validate.js schema, then semantics: edge ends exist, ids are unique,
one accent, groups known, acyclic, at most two deep.
EVERY error, each with a JSON Pointer, what was
expected and the valid options:
/edges/1/to: unknown node "billing". Known ids: in, jev, bill, rev.
diagram/repair.js the ladder. prepare() takes one call up it.
The ladder departs from the spec in one place, on purpose. The spec lists
"dedupe ids" and "drop edges to unknown nodes" among the fixes made without a
model, while its own table says semantic errors go to the model, "because
intent is needed". Both cannot hold: an edge to "billing" when the node is
"bill" is a typo the model fixes in one pass, and dropping it silently changes
what the diagram says. So:
1 auto-fix, lossless only lenient JSON (comments, trailing commas, smart
quotes, bare keys), synonyms ("diamond" for
decision, source/target for from/to), slugged
ids, an edge that names a node by its label,
long labels shortened with the full text kept
2 every error, once returned for the model's one repair pass
3 degrade only now the lossy fixes: edges to nowhere
dropped, duplicates renamed, one accent kept,
counts up to the HARD tier — each loss named
An auto-fixed diagram never says something different from what the model
wrote; a degraded one always says what it lost.
accept() is the same gate for a spec that is read back rather than received —
from a session file, say: validate, hand on only the fields the schema
declares, and fall back to the rungs that need no model.
No Ajv: it is a dependency, and this extension has neither dependencies nor a
build step.
test/fixtures/diagrams/corpus.json holds 39 broken specs with the exact outcome
of each attempt. They are seeded from the ways models are known to get this
wrong, not collected in the field; real ones should be added as they turn up.
diagramSchema (18 tests) and diagramRepair (65).
…house style
diagram/theme.js the style guide as numbers: title 15, node name 13
semibold, second lines and edge labels 11.5, nothing
under 10.5; radius 8, border 1.25, padding 12; groups
inset 16; the accent a low-opacity fill and a 2px border.
Colours are editor theme tokens, with two fixed palettes
for exported files.
diagram/layout.js layout(spec, opts) gives geometry; inspect(geometry) lists
whatever overlaps, overflows or leaves the frame.
The spec names ELK.js. This is not ELK: ELK is EPL-2.0 in a repository that is
otherwise MIT-clean, about 1.5 MB, and this extension has no build step. The
spec's own open question asks whether a simpler layered layout is needed as a
fallback; this is that layout, as the only one. layout() is the single entry
point, so ELK can replace it without touching anything else.
What it does: ranks by longest path; crossings reduced by sweeps; positions
across the flow solved as constraints and then lined up by medians; a port per
edge; a track per connector in each gap, so lines cross but never run along
each other; and each edge label placed BESIDE its line, scored against nodes,
lines and other labels. A flow asked for left-to-right that will not fit its
column is laid out top-to-bottom instead.
A group's name sits in the top-left of its frame — which, when the flow runs
down, is exactly where lines come in. No connector runs through it. The name
slides along the frame to the nearest clear stretch; if there is none and the
lines have under 72px to move, they are moved to the far side of the name (the
frame grows by that much, never past the column); otherwise the name stays and
is marked to be drawn over the line.
Measured over 6,000 random specs at the model-facing limits: no node overlaps
another, no connector crosses a node or an unbacked group name, nothing leaves
the frame. Three known limits, each bounded in the suite so it cannot quietly
get worse:
an edge label that falls back to a halo over a line 6.5% of specs
two connectors that swap lanes in one gap 0.3%
a group's name with a connector behind it 2.7% of specs with groups
None of the nine gallery diagrams (fixtures/diagrams/gallery.json) shows any of
the three, at its natural width or at 560, 420 and 320px. The twelve-node one
lays out in under a millisecond once warm; the slowest of the 6,000 took 7ms.
diagramLayout (23 tests, one of them a 1,500-spec fuzz).
diagram/scene.js geometry to a tree of elements, built from two allow-lists:
seven tags (svg g rect path text title style) and a set of
attributes with no href, no style, no id and no on*.
mount() turns it into DOM with createElementNS and
textContent; toSvg() into an escaped string for a file. A
label is never parsed as markup by either, and both refuse
anything off the lists rather than trust whoever built the
tree.
diagram/text.js the outline a screen reader is given (nodes, then edges, in
reading order); the one-line stub that stands in for a spec
once a conversation is compacted; Mermaid and source
export.
diagram/ascii.js for when a picture cannot be shown: the SAME layout on a
character grid, so it puts things where the picture would.
Plain ASCII; East Asian wide characters and emoji count as
two cells, combining marks as none.
A linked node carries data-lc-link = its NODE ID, never the path: whoever
handles the click looks the link up in its own copy of the spec.
A group's name that the layout could not keep clear of a connector is drawn
after the connectors, with the halo an edge label gets, so the line reads as
passing behind the word.
Snapshots of three gallery diagrams in both palettes (fixtures/diagrams/
snapshots): a change to the style, the layout or the painter shows up there as
a diff. diagramScene (23 tests) and diagramText (17).
Ask the agent how something is put together and it answered with boxes and
arrows made of characters, which break with the font, the theme and the width
of the panel. Now it has a tool. It describes the structure; the editor
validates it, lays it out and paints it in the chat, in the editor's theme,
with nodes that open the code they stand for.
Graph JSON only: phase 1 of the spec. Mermaid, Vega-Lite and raw SVG are not
built, and the rules tell the model to use a table for numbers and a numbered
list for sequences rather than a format the chat would show as source.
The model's side (diagram/tool.js)
render_diagram its input schema is the validator's own schema object
the rules when a diagram earns its place; never draw with characters;
asked for a diagram, draw it here rather than writing a file
of diagram source; the title states the takeaway; one
accent; split above twelve nodes; link nodes to workspace
files. One worked example, itself a valid spec.
the answer {"ok":true,"id":"d-1"}, or every error at once
About 520 tokens of tool and 450 of rules on every request from a client
that can render — constant for a session, so they sit in the cached prefix.
One conversation's diagrams (diagram/service.js)
A call that fails validation is sent back ONCE, with every error. A second
failure is drawn degraded, with a banner naming each loss and a Retry button
that the user presses, not the editor. Three bounces in one run and every
later call is final. A diagram still owed its repair when the run ends — the
model gave up, ran out of steps, was stopped — is settled from its last
attempt before agentDone, so no placeholder is left waiting. Arguments cut
off at the token limit are asked for again, never repaired.
In the chat (media/chat.html)
A placeholder from the moment the call starts, titled as soon as the title
has streamed; then the picture. The shared modules are inlined into the page
by diagram/bundle.js under the page's existing nonce: the content security
policy is unchanged, and the page still loads nothing. Text is measured in
the real font. A column too narrow for a left-to-right flow turns it
downward, and it turns back when the column widens. A full-size view with
zoom and pan. The "auto-fixed" badge and its details; the degraded banner;
the errors and the source when nothing can be drawn. Copy source, SVG, PNG,
open as Mermaid, insert into a Markdown file. The title is the accessible
name, and a text outline is there for screen readers. If painting fails, or
the modules never loaded, the same diagram is shown as text — drawn by the
page in the first case and by the host in the second.
What the page may ask of the editor
openLink, export, retry and ascii. Each names a diagram by id and is looked
up in the host's own records. A click sends the NODE id; the path comes from
the record and is resolved again at that moment (diagram/links.js): inside a
workspace folder once ".." and symlinks are followed, and a file. SVG and PNG
come back from the page, so they are checked before they are written
(diagram/exportCheck.js). The spec lists two actions; its own Retry button
and text fallback need the other two.
Sessions and context
The final record — spec, status, fixes, what was lost — is stored with the
turn and replayed when a chat is reopened; nothing is repaired again. A
session file is input too: a stored spec passes the validator on its way
back in, in the host and in the page. At compaction a diagram becomes one
line ("diagram: <title>, 4 nodes, id d-17"), and get_diagram is offered from
then on. A session export writes each diagram as a fenced Mermaid block.
Who gets it
A run carries client.render: rich or ascii. Rich is the chat webview, unless
levelcode.ai.diagrams.enabled is off or the model's catalog row says
diagrams: false — the switch for a model that fails the eval. An ascii client
is offered neither the tool nor the rules. Chat mode has no tools and is
unchanged.
Counters (diagram/stats.js, "AI: Diagram Statistics")
First-pass valid rate, fixes by rung, error classes per model, render time,
tokens per diagram, and answers that drew with characters anyway. Kept in
the editor's own storage and sent nowhere; enums and numbers only, never a
label, a title or a path.
Providers: onToolStart now carries the call's id, a new onToolInput streams its
arguments, and a turn returns the raw text of arguments that did not parse.
All additive.
Four existing suites change. agentNoWorkspace and sessionsUi pinned the shape
of the tool list in source, and now pin the new shape. authRetryCallers and
sessionExpiredHost slice compactAgentMemory and resumeSession out of
extension.js, and needed the names those now use.
Tests: diagramAgent (22 — the real runAgent with a scripted provider),
diagramSession (12), diagramLinks (9), diagramHost (27 — extension.js's own
functions against a stand-in for vscode), diagramStats (8), diagramUi (11).
…, and an eval for models Three scripts for the three things the suites cannot show. None is in the gate, which is plain Node and stays that way. scripts/diagram-browser-check.js — the shipped chat.html in headless Chrome, built the way the host builds it (the same bundle, the same policy) and fed records made by the real service. 155 checks across both themes: every diagram is an SVG of the painter's elements only; hostile labels are on the page as text and nothing ran; no policy violation and no request; the placeholder, the badge, the banner, the failed view; a click names a node, never a path; SVG and PNG export pass the host's own check; and a painter that throws, or a page with no diagram modules at all, both end in the text fallback. Each step waits for its result, not for a length of time: the first version waited, and failed one run in three. scripts/diagram-editor-check.js — the real editor. A second, throwaway instance of the dev build (its own profile, sessions folder and workspace) loads THIS checkout's extension in place of the built-in one, talks to a stand-in provider on localhost, and is driven through the DevTools protocol. 17 checks: the tool and its rules reach the model; the diagram is painted in the real webview in the editor's colours; a wider column re-lays it out; a linked node opens its file at the symbol; the session on disk holds the diagram; and the answer around it, which quotes a code fence, is one paragraph (the chat fix earlier in this branch). It exists because of a mistake worth writing down. The feature was first called done having only ever run in a browser, with instructions for trying it that could not have worked: the work sat uncommitted in a git worktree, and a worktree has no vscode/ — run-dev.sh runs the extensions of the checkout that does. The editor that was then tried was develop's, with no diagrams in it. What does work, by hand, from the checkout that has vscode/: ./scripts/run-dev.sh --extensionDevelopmentPath=<checkout>/extensions/levelcode-ai scripts/diagram-eval.js — the spec's eval: 30 prompts that should produce a diagram and 10 that should not (fixtures/diagrams/eval-prompts.json), each run through the real agent loop and scored on format choice, first-pass validity, fixes by rung, degraded rate, error classes, tokens, node-count overruns, title quality and answers that drew with characters. --dry-run uses a scripted model and no network. --run makes billed calls on your key, says how many before the first one, and is refused without a model and a key; the workspace is an empty temporary folder and every approval is answered no. It has not been run against any model. That is the next thing this feature needs: nothing here says how often a real model's first attempt is valid, or whether it reaches for the tool at all. diagramEval (16 tests) holds the harness to its own arithmetic: a scripted model whose mistakes are known comes out at exactly the numbers its script implies, and nothing is sent without --run, a model and a key.
…ow to check it docs/RICH-DIAGRAMS.md is the implementation record for the Rich Diagrams spec (2026-10-04), under the spec's own section names, because the code cites them. Per section: the rule, what the code does about it, and each place the build chose differently, with the reason, so the choice can be argued again. A status per requirement; the known limits with their measured rates; what is verified and what is not — no live model has drawn a diagram yet, and the eval has not been run. It is not the spec. The spec is a private document; its links, and one line about pricing, are left out of a public repository. CLAUDE.md gains the feature's entry and three conventions that cost time to learn: the shared diagram modules are pasted into a script block and must never spell a script tag or an HTML comment opener; the host suites' brace matcher cannot read a backtick inside a regex literal; and work in a worktree is not in the editor until it is loaded with --extensionDevelopmentPath. test/fixtures/diagrams/README.md says what each fixture is for, and how to add a real broken spec to the corpus.
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Untrusted schema versions can crash validation, and several repair, replay, compaction, statistics, and accessibility paths remain incorrect.
Review effort: Balanced
Findings: 2
Open (7)
Safely handle externally supplied model IDs · New Validate schema versions against own registry entries · New Preserve parsed identity during lenient repair · New Exclude records superseded by replacement records · New Trap focus within the modal dialog · New Avoid nested interactive elements in diagram stage · New Preserve text and diagram block order · New
What changed in this PR
Adds phase-one rich Graph JSON diagrams to agent chat, including validation, rendering, persistence, export, accessibility fallbacks, and evaluation tooling. It also fixes streamed Markdown fence handling.
Changes:
- Adds the
render_diagramagent tool and provider streaming support. - Implements safe diagram layout, rendering, links, exports, persistence, and statistics.
- Adds extensive unit, browser, editor, snapshot, and evaluation coverage.
| File | Description |
|---|---|
CLAUDE.md |
Documents diagram architecture and conventions. |
docs/RICH-DIAGRAMS.md |
Records the feature design. |
extensions/levelcode-ai/agent.js |
Integrates diagram tools into agent runs. |
extensions/levelcode-ai/diagram/ascii.js |
Adds text fallback rendering. |
extensions/levelcode-ai/diagram/bundle.js |
Bundles shared modules into the webview. |
extensions/levelcode-ai/diagram/exportCheck.js |
Validates exported SVG and PNG data. |
extensions/levelcode-ai/diagram/layout.js |
Implements layered diagram layout. |
extensions/levelcode-ai/diagram/links.js |
Resolves safe workspace links. |
extensions/levelcode-ai/diagram/repair.js |
Implements normalization and repair. |
extensions/levelcode-ai/diagram/scene.js |
Builds allow-listed SVG scenes. |
extensions/levelcode-ai/diagram/schema.js |
Defines the versioned Graph JSON schema. |
extensions/levelcode-ai/diagram/service.js |
Manages diagram lifecycle and records. |
extensions/levelcode-ai/diagram/stats.js |
Collects local rollout metrics. |
extensions/levelcode-ai/diagram/text.js |
Generates outlines, stubs, and Mermaid. |
extensions/levelcode-ai/diagram/theme.js |
Defines diagram styling and theme tokens. |
extensions/levelcode-ai/diagram/tool.js |
Defines tools and model instructions. |
extensions/levelcode-ai/diagram/validate.js |
Validates schema and semantics. |
extensions/levelcode-ai/extension.js |
Wires host actions, persistence, and exports. |
extensions/levelcode-ai/media/chat.html |
Adds diagram cards and zoom UI. |
extensions/levelcode-ai/package.json |
Registers settings and statistics command. |
extensions/levelcode-ai/providers/anthropic.js |
Streams tool IDs and partial arguments. |
extensions/levelcode-ai/providers/catalog.js |
Adds per-model diagram capability. |
extensions/levelcode-ai/providers/index.js |
Forwards diagram streaming callbacks. |
extensions/levelcode-ai/providers/openaiCompat.js |
Handles streamed OpenAI tool arguments. |
extensions/levelcode-ai/providers/translate.js |
Preserves malformed raw tool arguments. |
extensions/levelcode-ai/scripts/diagram-browser-check.js |
Adds browser-level rendering checks. |
extensions/levelcode-ai/scripts/diagram-editor-check.js |
Adds real-editor integration checks. |
extensions/levelcode-ai/scripts/diagram-eval.js |
Adds model evaluation harness. |
extensions/levelcode-ai/sessionEvents.js |
Stores and exports diagram events. |
extensions/levelcode-ai/sessions.js |
Persists diagrams with sessions. |
extensions/levelcode-ai/test/agentNoWorkspace.test.js |
Updates host-gated tool assertions. |
extensions/levelcode-ai/test/authRetryCallers.test.js |
Stubs diagram compaction behavior. |
extensions/levelcode-ai/test/chatMarkdownFences.test.js |
Tests streamed Markdown fence handling. |
extensions/levelcode-ai/test/diagramAgent.test.js |
Tests agent diagram integration. |
extensions/levelcode-ai/test/diagramEval.test.js |
Tests evaluation logic. |
extensions/levelcode-ai/test/diagramHost.test.js |
Tests host wiring and security. |
extensions/levelcode-ai/test/diagramLayout.test.js |
Tests layout and fuzz cases. |
extensions/levelcode-ai/test/diagramLinks.test.js |
Tests workspace-link containment. |
extensions/levelcode-ai/test/diagramRepair.test.js |
Tests the repair ladder. |
extensions/levelcode-ai/test/diagramScene.test.js |
Tests SVG safety and snapshots. |
extensions/levelcode-ai/test/diagramSchema.test.js |
Tests schema validation. |
extensions/levelcode-ai/test/diagramSession.test.js |
Tests persistence and replay. |
extensions/levelcode-ai/test/diagramStats.test.js |
Tests metric aggregation and privacy. |
extensions/levelcode-ai/test/diagramText.test.js |
Tests textual representations. |
extensions/levelcode-ai/test/diagramUi.test.js |
Tests webview integration and safety. |
extensions/levelcode-ai/test/fixtures/diagrams/README.md |
Documents diagram fixtures. |
extensions/levelcode-ai/test/fixtures/diagrams/corpus.json |
Adds malformed-spec corpus. |
extensions/levelcode-ai/test/fixtures/diagrams/eval-prompts.json |
Adds model evaluation prompts. |
extensions/levelcode-ai/test/fixtures/diagrams/gallery.json |
Adds representative diagrams. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/arch.dark.svg |
Adds dark architecture snapshot. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/arch.light.svg |
Adds light architecture snapshot. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/decision.dark.svg |
Adds dark decision snapshot. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/decision.light.svg |
Adds light decision snapshot. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/jev.dark.svg |
Adds dark routing snapshot. |
extensions/levelcode-ai/test/fixtures/diagrams/snapshots/jev.light.svg |
Adds light routing snapshot. |
extensions/levelcode-ai/test/sessionExpiredHost.test.js |
Adds diagram service test stub. |
extensions/levelcode-ai/test/sessionsUi.test.js |
Updates session integration assertions. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…eant to be; a dead store goes Two findings from the code-quality review. Both are right. theme.js, charEm() — "the condition c > 0xffff is always false". The width estimate used where there is no canvas asked "CJK and other full-width scripts" (U+2E80 and up, 1.0 em) before "emoji / astral" (above U+FFFF, 1.1 em). Every code point above U+FFFF is above U+2E80 too, so the second question was never reached and an emoji was measured at 1.0 em. The estimate is meant to err wide — the failure that matters is text overflowing its box — so the questions are now asked the other way round. Nothing on screen changes. The chat measures with the real font, and no fixture contains a character beyond the basic plane, so no snapshot moves. What changes is the fallback when the canvas refuses the font, and the estimates made in Node. layout.js — "the value assigned to span here is unused". When several connectors share a side too short for them, the node grows to hold exactly what they need; span was recomputed afterwards and never read. The recompute is gone and span is a const. it.cs, set on the same line, IS read later, and stays. 6,000 random layouts and the nine gallery diagrams are byte-identical to the code before. diagramLayout gains a test for the estimate's classes: narrow, ordinary and wide letters, digits and capitals, a CJK character at a full em, an emoji at 1.1 em and counted as one character. It fails on the old order, on exactly the condition the review named. 62 suite files pass on macOS and in a Linux container.
…model, a version or a shape
Two findings from the Copilot review, both right, and a third of the same kind
that looking for them turned up.
stats.js — "IDs such as `constructor` resolve inherited properties". The
counters are plain objects keyed by model id, and a model id is whatever a
provider calls it. For a model called "constructor" the lookup found Object
itself: `calls++` wrote to Object, and the next line threw. For "__proto__" it
found the prototype every object shares, and `calls++` wrote there — every
object in the extension host had a `calls` of NaN — before it threw. The service
swallows a counter that throws, so the diagram was still drawn; the write was
not undone. An error class is a key too: "constructor" passes the class
pattern, and its count came out as "function Object() { [native code] }1".
A key is now read as an own property and written as one (defineProperty, so
"__proto__" is an entry and not a call to the prototype's setter). The store
stays an ordinary object: what is already in the editor's storage reads as
before, and it survives the JSON round trip it is stored through.
validate.js — "a crafted schema version such as `toString` resolves through
SCHEMAS' prototype". validate() promises a list of errors and threw instead:
SCHEMAS["toString"] is a function, so the version was "known", and the check
that followed read `.type` of undefined. prepare() threw with it, and so did
accept(), which is what reads a stored record back. A version is now looked up
only when it is one of the numbers this editor reads. That is deliberately
stricter than "an own key of the registry": no schema declares `v`, so the
string "1" passed as version 1 and was stored as a string, with nothing after
the lookup to object. On the ladder a version written as digits is still read
as the number, as before.
repair.js — the same lookup a third time, found by feeding those names through
every field. The synonym tables are objects, and `"shape": "constructor"`
found Object: the spec came out with a FUNCTION for a shape, and the model was
told "/nodes/0/shape: expected a string, got function" about a string it had
written. A line style of that name was read as solid without a word. A word is
now looked up as the table's own entry, so this one is unknown like any other
unknown word: drawn with the default, and said so.
Nothing else in these modules indexes an object with text from a spec: ids,
labels and groups are kept in Maps, and the schema is walked by its own field
names. The chat page keeps a table of its own, keyed by tool-call id; that one
is fixed with the page, two commits on.
Each has a test that fails on the code as reviewed: six such names as model
ids and one as an error class, with nothing written to Object or its
prototype, before and after a trip through JSON; fifteen things that are not
version 1 — the string "1" and the list [1] among them — as one error each and
never a throw, and four of them up the ladder and through accept(); seven such
words as a direction, a shape and a line style, with exactly the outcome of
"zigzag".
…laced drawing is not handed back, and a replay keeps the order
Three findings from the Copilot review. All three are right.
agent.js:1041 — "this discards parsed.value and passes the raw string … the
original is degraded and a second card is drawn". The service decides whether a
call is the waiting diagram's repair by its title and its node ids, and it read
them off the arguments as they arrived. Arguments that are loose JSON arrive as
text — no title, no nodes — so a repair written with a trailing comma was a
"different diagram": the first attempt was settled as a degraded picture of its
own and the correction was drawn as a second. The same went for a first attempt
in loose JSON (its repair never found it), for a spec inside a `{"spec": …}`
envelope, and for one that says `name` for `title`.
The identity is now read from the spec as the ladder reads it — normalize(),
the lossless rung — whatever form it came in. agent.js is unchanged: it still
hands over the text, so the loose JSON is still counted as the auto-fix it was.
The placeholder's title comes from the same reading, so a call that arrived as
text has one too. And a second drawing of the same title in one run is
recognised as a redraw when it arrives as text.
service.js:260 — "this emits stubs for records that a later drawing replaced".
A redraw replaces the earlier drawing in the chat and a reopened session leaves
it out, but at compaction the model was handed a stub for each — so it could
see, fetch and edit a diagram the user no longer has. A replaced drawing now
gets no stub, wherever its replacement is in the conversation; it is not among
the known ids; and get_diagram, asked for it by an id the model may still
remember, answers with the id of the drawing the chat shows — the end of the
chain when it was redrawn more than once. A record that names itself, or two
that name each other, is a damaged file and not a reason to hang.
sessionEvents.js:149 — "all text blocks are collapsed and emitted before the
diagram". A message that says something, draws, and says more was replayed as
all of its prose and then the picture. It is now walked block by block: prose,
picture, prose, as it was on screen. Every piece after the first carries
`cont: true`, so a reader of the turns can still tell where a message ends
(the chat page does not need it: it already gives a run of answers one
speaker label). A message with no picture in it comes out exactly as it did,
one turn of all its text — and so does every caller that passes no records.
The Markdown export is built on those turns, and getting its order right
showed two more things wrong with it. A picture with no prose before it was
written under whoever spoke last — after a question, under "You". And once a
message is split, the prose after its picture would have been a turn of its
own. The export now works message by message: one rule and one speaker label
per message that says something, its pictures where they stood; a message that
is only a picture joins the answer above it, as before, unless there is none;
a diagram that was never drawn is dropped only after the pieces are grouped,
so the prose after it is not appended to the message before; and the count in
the header is still the number of messages that said something.
Tests, each failing on the code as reviewed. Through the real agent loop: a
repair in loose JSON, a first attempt in loose JSON, three envelope and
renamed-field shapes, a repair with the title reworded, and a redraw in loose
JSON — one diagram on file each time — and a different diagram in loose JSON,
which must NOT be taken for the same one. Three drawings of one diagram leave
one stub, and none in a stretch that holds only a replaced one; the same after
a reload, and in a file whose records name each other. A message of prose,
picture, another tool's call, prose and picture replays in that order with the
same words, and is exported as expected in five shapes.
… is not a button, and one diagram is one card
Two findings from the Copilot review, both right. Checking them in a browser,
with the messages in the order a real run sends them, showed three more faults
in the same code: the page's bookkeeping of which card a diagram is drawn in.
chat.html:1523 — "this is declared modal, but opening it neither makes the
background inert nor traps Tab focus". The full-size view said
aria-modal="true" and covered the page, and the chat behind it could still be
tabbed into. While it is open, every other child of the body is now inert —
not focusable, not clickable, not read out — and released when it closes,
before the focus is handed back (an inert element cannot take it). Tab goes
round inside: past the last stop is the first, before the first is the last,
and a focus that got outside is brought back. A linked node in the view could
be tabbed to and did nothing on Enter; it opens its file now, as a click does.
chat.html:4442 — "the stage is exposed as a button while linked SVG nodes
inside it are independently focusable links". A button's content is no place
for other controls. The picture is no longer a control: no role, no tab stop,
no name. Opening it full size is a button of its own, first in the toolbar
("Full size"), and that is what the keyboard and a screen reader use; a click
on the picture still does the same, as a shortcut for the pointer. Where there
is no picture — the text fallback — there is no such button.
The three it turned up. agent.js announces every render_diagram call the
moment it starts streaming, before anyone can know whether it is a new
diagram, a repair or a redraw, and the page's checks had never been played in
that order.
- A redraw left BOTH drawings on screen, where a reopened session shows one.
The new call had made a placeholder of its own, the record was drawn in it,
and the drawing it replaces stayed. The earlier one now goes when its
replacement arrives, and the redraw stands where the model drew it again —
whether or not the call announced itself.
- A different diagram arriving while one was being repaired painted over it.
The page lets the next call take over a card that is waiting for a repair,
which is right when it is the repair. When it is not, the host settles the
waiting diagram into that card, and the newcomer then drew over it: two
diagrams on file, one on screen. A record now has a claim on a card that is
still a placeholder, or that shows the drawing it replaces, and on no other.
- A tool call named `__proto__` broke the ones after it. The table of cards
was a plain object keyed by tool-call id, a name a provider chooses.
`__proto__` made that card the table's prototype; the next lookup of
`parentNode` threw "Illegal invocation", and that diagram was never shown.
The table has no prototype now.
Tests. The gate cannot boot the page, and until now it only pinned its source;
every pin was green through all five faults. Two parts of the page are now RUN
there, sliced out of the file. The card bookkeeping, against a stand-in for
the few DOM calls it makes: a repair, two calls in one turn, a redraw with and
without a placeholder of its own, a change of subject, a record with no claim
on the card it points at, seven tool-call ids that are also names the table
could inherit, a wiped transcript, and the sweep at the end of a run — each
judged by what is on screen and where. And the function that decides where Tab
goes.
The browser check gains the modal's behaviour in real Chrome — inert, a focus
that cannot leave, Tab at both ends, Enter on a node, Escape and the focus back
on the button — and a page that plays seven runs in the real order, then
replays each as a reopened session and compares the two. 186 checks, up from
155. The real-editor check opens the view from its button in the real webview
and confirms the page behind it is inert there: 19 checks, up from 17.
…ng only the chat can open Not from the review: found while fixing what it said about the picture and the linked nodes inside it. A linked node is a focusable `role="link"` group carrying `data-lc-link`. In the chat, a click or Enter asks the host to open the file. "Save as SVG" wrote the same attributes into the file, where nothing can answer them: opened in a browser, the picture had tab stops that announced "Open agent.js at runAgent, link" and did nothing. A standalone build — the one made for a file — now leaves the four attributes off. The node keeps what a picture can use: its file icon, and the tooltip with the path. The picture in the chat and its full-size view are unchanged. The two arch snapshots lose exactly those attributes on their two linked nodes and nothing else; the other four snapshots do not move. diagramScene walks the standalone tree for all four, and checks both nodes are still marked as files.
RICH-DIAGRAMS.md after the #104 review round. What it now says about the feature: which call counts as a diagram's repair (its title or its nodes, as the ladder reads them); that a reopened chat shows what the live one showed — prose and pictures in the order they were written, one diagram one card, a replaced drawing gone from both; that a replaced drawing is not handed back to the model; that the full-size view is modal and is opened by a button of its own; that a saved SVG carries no links; and that a name arriving from outside is never used to index a plain object by itself. One new known limit: a repaired diagram is drawn live where its first attempt was waiting, and a reopened session shows it at the call that got it right. And the numbers, as run on this tree: 270 tests in the twelve diagram suites (from 252), 186 checks in the browser (from 155), 19 in the real editor (from 17), the gate on macOS and in the Linux container, the suites with every timer delayed, and 69 single-edit mutations against this round's changes, all caught after two gaps were closed. The first mutation set was not run again, and the doc says so. The count of suite files in the whole gate is gone from the sentence about it: it goes stale with every branch that adds a test file.
ndemianc
added a commit
that referenced
this pull request
Oct 8, 2026
One conflict, in CLAUDE.md: develop (#104) added bullets around the "Commit ..." line that this branch edits to list `modules/`. Both are kept — the merged file is develop's plus exactly this branch's changes. The gate (scripts/test-extensions.sh) passes on the merged tree: 64 test files, develop's diagram suites and this branch's module suite included.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


What this is
Ask the agent how something is put together and it answers with boxes and arrows typed out of characters. They break with the font, the theme and the width of the panel; they cannot be clicked or exported.
With this PR the agent has a tool,
render_diagram. The model describes structure — nodes, edges, groups, one accent, never a coordinate or a colour. The editor validates the description, lays it out in one house style and paints it in the chat, in the editor's theme, with nodes that open the code they stand for.This is phase 1 of the Rich Diagrams spec: Graph JSON only. Mermaid, Vega-Lite and raw SVG are not built.
Two of the committed snapshots: what the SVG export writes. In the chat the colours are the editor's own theme tokens.
How it works
diagram/schema.jsvalidate.jsrepair.jsdiagram/theme.jslayout.jsdiagram/scene.jstext.jsascii.jstextContent), the screen-reader outline, Mermaid and source export, the text fallbackdiagram/tool.jsservice.jsdiagram/links.jsexportCheck.jsstats.jsbundle.jsagent.js,providers/*sessionEvents.jssessions.jsextension.jsget_diagram; link, export, retrymedia/chat.htmlThe page's content security policy is unchanged. The shared modules are inlined under the nonce the page already has, and it still loads nothing.
Where it departs from the spec
layout.layout()is the one entry point, so ELK can replace itbillingwhen the node isbillis a typo; dropping it silently changes what the diagram saysopenLink,export,retry,asciiThe rule the ladder follows: an auto-fixed diagram never says something different from what the model wrote; a degraded one always says what it lost.
Also in here: a chat rendering bug, unrelated to diagrams
The first two commits fix a bug that is on
developtoday. An answer that mentioned a code fence in passing — the model quoted three backticks inside an inline span of four — was rendered as two one-character code blocks, and the rest of the paragraph came out one streamed fragment per line.The renderer treated any three backticks, anywhere, as a fence. And once a "closing" fence had no newline after it, the streaming renderer froze every delta into its own block. A fence is now a line of its own, inline code can be delimited by any number of backticks, and only a block whose closing line is complete is frozen.
96399b6pins what the renderer did before;3b05fd3is the fix. Replayed in Chromium with the session's exact text in 184 deltas: the paragraph was 57 lines and two code blocks before, and is one paragraph and none after. The two commits stand alone, and can be cherry-picked into their own PR if they should merge first.The commits
96399b63b05fd342159e5fb2e67af35fbc1ce02534render_diagramin the agent, the host and the chat822cfc2db68f68docs/RICH-DIAGRAMS.md,CLAUDE.md, the fixtures' READMEa317a45e84b36849d16b851e44900f551873c51f11Each one passes
./scripts/test-extensions.shon its own. The gate was run on every committed tree in a separate checkout, not only on the tip.Review
Nine findings in two rounds: two from the code-quality check and seven from Copilot. All nine were right. Each is fixed with a test that fails on the code as reviewed, and each thread is answered and resolved.
theme.js:c > 0xffffis always falsea317a45)layout.js: a value assigned and never reada317a45)stats.js: a model calledconstructoror__proto__Object, or the prototype every object shares, and wrote a counter to it. Keys are read and written as own entries (e84b368)validate.js: a schema version calledtoStringvalidate,prepareandacceptthrew where they promise an error. A version must be one of the numbers this editor reads, which also stops the string"1"passing as version 1 (e84b368)agent.js: a repair written as loose JSON loses its identity49d16b8)service.js: stubs for drawings that were replacedget_diagramon it names the drawing the chat shows (49d16b8)sessionEvents.js: prose collapsed ahead of the diagram49d16b8)chat.html: the full-size view is modal in name only51e4490)chat.html: the picture is a button with links inside it51e4490)Fixing these turned up six faults nobody had reported. They are fixed in the same commits, and the last has one of its own.
__proto__made the page throw, and a later diagram was never shown."shape": "constructor"foundObjectin the synonym table. The model was told "expected a string, got function" about a string it had written.0f55187).The first three are in the chat page, and neither net could have caught them. The gate only pinned the page's source, and the browser check never played messages in the order a real run sends them: a placeholder at the start of every call, before anyone knows whether it is a new diagram, a repair or a redraw. Both gaps are closed. The page's card bookkeeping now runs in the gate, and the browser check plays seven runs in the real order, then replays each as a reopened session and compares the two.
Verification
The numbers are for the tip,
3c51f11.developatedc2752in a scratch checkout, 63 pass.Extension unit testsand the CodeQL analysis pass on GitHub's runner for3c51f11.runAgentloop against a scripted provider,extension.js's own functions sliced out and run against a stand-in forvscode, and the page's card bookkeeping run against a stand-in for the few DOM calls it makes.scripts/diagram-browser-check.js): the shippedchat.htmlin headless Chrome under the real policy, 186 checks. Hostile labels are on the page as text and nothing ran; no policy violation; no request. The full-size view's focus is checked in real Chrome, and seven runs are played in the real order, live and then reopened.scripts/diagram-editor-check.js): a second, throwaway instance of the dev build loads this branch's extension and talks to a stand-in provider on localhost, 19 checks. The tool and its rules reach the model, the diagram is painted in the real webview in the editor's colours, the full-size view opens from its button and is modal there, a linked node opens its file at the symbol, and the session on disk holds the diagram.Not verified
scripts/diagram-eval.jsis the spec's eval — 30 prompts that should produce a diagram and 10 that should not, run through the real agent loop — and it has not been run, because--runmakes billed calls. Nothing here says how often a real model's first attempt is valid, or whether it reaches for the tool at all.inert, in Chrome and in the editor's own webview. Nobody listened to one.For the reviewer to decide
levelcode.ai.diagrams.enableddefaults to true, and costs about 970 tokens of tool and rules on every agent request while it is on. It is constant for a session, so it sits in the cached prefix. Flip the default if this should ship dark until the eval has run.diagrams: falsein itsproviders/catalog.jsrow. That is the lever for one that fails the eval.docs/RICH-DIAGRAMS.mdis an implementation record, not the spec. The spec is a private document.Known limits
Measured over 6,000 random specs at the model-facing limits. None of the nine gallery diagrams shows any of the three, at its natural width or at 560, 420 and 320 px.
Each is bounded in the layout suite, so it cannot get worse quietly. The rest are in the doc.
One more, from the review round: while a run is live, a repaired diagram is drawn where its first attempt was waiting, and a reopened session shows it at the call that got it right. The cards are the same; the place differs by whatever the model wrote in between.
To try it
On this branch, in the checkout that has
vscode/,./scripts/run-dev.shis enough. From a git worktree, which has novscode/of its own, load the worktree's extension into the build that exists:The title bar says
[Extension Development Host]. In agent mode, ask something whose answer is structure: "how does a request flow through this app?"